Papers with misaligned tokenization methods

1 papers
Jeff Da at COIN - Shared Task: BIG MOOD: Relating Transformers to Explicit Commonsense Knowledge (D19-60)

Copied to clipboard

Challenge: Recent studies show that large-scale pre-training models can be effective for large datasets.
Approach: They propose a method of integrating contextual embeddings with commonsense graph embeddINGs by preprocessing knowledge bases and aligning tokens between misaligned tokenization methods.
Outcome: The proposed method achieves higher accuracy than BERT and scores highest without pretraining.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations